Papers with multi-modal Transformer

2 papers
Visio-Linguistic Brain Encoding (2022.coling-1)

Copied to clipboard

Challenge: Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning.
Approach: They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning .
Outcome: The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing .
Encoding and Controlling Global Semantics for Long-form Video Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to find answers for long videos fail to reason over the whole sequence of video, leading to sub-optimal performance.
Approach: They propose a state space layer to integrate global semantics into video . they use a gating unit to enable controllability over the flow of global semantic into visual representations.
Outcome: The proposed framework is able to integrate global semantics into visual representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations